Papers with annotation procedure
Decontextualization: Making Sentences Stand-Alone (2021.tacl-1)
Copied to clipboard
| Challenge: | Taking excerpts of text can be problematic, as key pieces may not be explicit in a local window. |
| Approach: | They define a problem of sentence decontextualization by rewriting a sentence to be interpretable out of context while preserving its meaning. |
| Outcome: | The proposed method can be used in question answering and document understanding tasks. |
NorDiaChange: Diachronic Semantic Change Dataset for Norwegian (2022.lrec-1)
Copied to clipboard
| Challenge: | NorDiaChange is the first dataset of diachronic semantic change on the lexical level for Norwegian. |
| Approach: | They describe a manual annotation process for a new dataset of diachronic semantic change for Norwegian. |
| Outcome: | The proposed dataset covers the time periods related to pre- and post-war events, oil and gas discovery in Norway, and technological developments. |
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)
Copied to clipboard
| Challenge: | Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications. |
| Approach: | They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages. |
| Outcome: | The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages. |
Designing a Russian Idiom-Annotated Corpus (L18-1)
Copied to clipboard
| Challenge: | a pilot experiment using the idiom-annotated corpus of Russian is described . corpora that could be used for training idiomatic classifiers are scarce, especially if one turns to other languages. |
| Approach: | They describe the development of an idiom-annotated corpus of Russian . the corpus is compiled from freely available online resources . |
| Outcome: | The proposed corpus is based on an online corpus of Russian texts . it is available for research purposes and can be used for linguistic studies and pedagogy . |
Annotation of metaphorical expressions in the Basic Corpus of Polish Metaphors (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Polish texts annotated with metaphorical expressions is composed of two parts of comparable size, selected from two subcorpora of the Polish National Corpus . |
| Approach: | They propose to use a procedure to annotate metaphorical expressions in Polish texts using two different subcorpora of the Polish National Corpus . they propose several features to classify metaphorical Expressions identified in texts. |
| Outcome: | The proposed procedure is based on the MIPVU procedure and focuses on neologistic derivatives that have metaphorical properties. |
ODIL_Syntax: a Free Spontaneous Spoken French Treebank Annotated with Constituent Trees (2020.lrec-1)
Copied to clipboard
| Challenge: | ODIL Syntax is a French treebank built on spontaneous speech transcripts . the structure of every speech turn is represented by constituent trees . |
| Approach: | They propose a French treebank built on spontaneous speech transcripts with a constituency tree representation. |
| Outcome: | The proposed treebank is based on the French TreeBank, with some annotation guidelines . the proposed tree bank will be freely distributed by January 2020 under a Creative Commons licence . |
Budget Argument Mining Dataset Using Japanese Minutes from the National Diet and Local Assemblies (2022.lrec-1)
Copied to clipboard
| Challenge: | Budget argument mining attempts to identify argumentative components related to a budget item . argument mining is a subtask of QA Lab-PoliInfo-3 in NTCIR-16 . |
| Approach: | They propose a dataset to link budget information to minutes and budget items . they describe the construction of the dataset and the annotation procedure . |
| Outcome: | The proposed dataset is based on a QA Lab-PoliInfo-3 subtask . it identifies argumentative components related to a budget item and classifies them . |
M-CNER: A Corpus for Chinese Named Entity Recognition in Multi-Domains (L18-1)
Copied to clipboard
| Challenge: | NER is one of the most important natural language processing tasks. |
| Approach: | They propose to annotate sentences from human-computer interaction, social media, and e-commerce using two rounds of annotation. |
| Outcome: | The proposed system performs the best on all the data sets. |
Decomposing Unitization and Typing for Efficient and Consistent Span-Bound Concept Annotation (2026.findings-acl)
Copied to clipboard
| Challenge: | Substantial resources are typically spent on unitizing, the task of identifying precise span boundaries for entity mentions. |
| Approach: | They propose a method that focuses manual efforts on typed position annotations instead of full concept annotation. |
| Outcome: | The proposed procedure reduces the cost of concept annotations by focusing on typed positions instead of full concept annotation. |